Papers with linguistic theory
What Can String Probability Tell Us About Grammaticality? (2026.tacl-1)
Copied to clipboard
| Challenge: | linguistic theories have argued that language models have largely achieved grammatical competence, but they will assign non-zero probability to all strings. |
| Approach: | They propose a theoretical framework for analyzing string probabilities in linguistics based on simple assumptions about the generative process of corpus data. |
| Outcome: | The proposed framework makes three predictions using 280K sentence pairs in English and Chinese. |
Joint Universal Syntactic and Semantic Parsing (2021.tacl-1)
Copied to clipboard
| Challenge: | Several attempts have been made to jointly parse syntax and semantics, but this trade-off is not well understood. |
| Approach: | They propose multiple model architectures that exploit the rich syntactic and semantic annotations contained in the Universal Decompositional Semantics dataset to obtain state-of-the-art results. |
| Outcome: | The proposed model outperforms existing models in 8 languages and their results are consistent across languages. |
The Importance of Modeling Social Factors of Language: Theory and Practice (2021.naacl-main)
Copied to clipboard
| Challenge: | Current NLP models focus on information content while ignoring language’s social factors. |
| Approach: | They propose that NLP systems focus on information content while ignoring language’s social factors to improve performance. |
| Outcome: | The proposed approach improves the performance of existing systems, open up new applications, and increase fairness and usability for all users. |
Linguistically Grounded Analysis of Language Models using Shapley Head Values (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for probing language models for morphosyntactic constructions are not well understood . language models gain knowledge of grammatical phenomena during pretraining, but exactly how this knowledge is encoded is not well established. |
| Approach: | They propose a method for probing language models via Shapley Head Values . they use a BLiMP dataset to test their method on linguistic constructions based on a Shaply Head Value method . |
| Outcome: | The proposed method can be used to investigate linguistic knowledge in language models . it shows that attention heads responsible for processing related linguistic phenomena cluster together . |
Dead or Murdered? Predicting Responsibility Perception in Femicide News Reports (2022.aacl-main)
Copied to clipboard
| Challenge: | linguistic expressions of gender-based violence can conceptualize the same event from different perspectives by emphasizing certain participants over others. |
| Approach: | They conduct a large-scale perception survey of GBV descriptions from italian newspapers and train regression models that predict the salience of GV participants with respect to different dimensions of perceived responsibility. |
| Outcome: | The proposed model shows that salient focus is more predictable than salient blame, and perpetrators’ salience is more predictable than victims’ salient. |
Abstract Meaning Representation for Multi-Document Summarization (C18-1)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is a semantic representation of natural language based on linguistic theory . |
| Approach: | They propose to use Abstract Meaning Representation (AMR) as a content representation. |
| Outcome: | The proposed framework is fully data-driven and flexible. |
Debiasing Word Embeddings with Nonlinear Geometry (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods for debiasing word embeddings are limited to individual social categories . however, real-world corpora typically present multiple social categories that may correlate or intersect with each other. |
| Approach: | They propose a method to debias word embeddings using nonlinear geometry of individual biases. |
| Outcome: | Empirical results show that the proposed method mitigates biases associated with individual social categories and treats each category in isolation. |
No Questions are Stupid, but some are Poorly Posed: Understanding Poorly-Posed Information-Seeking Questions (2025.acl-long)
Copied to clipboard
| Challenge: | When a question is poorly posed, answerers struggle to converge on dominant interpretations, while models attempt comprehensive coverage by addressing many interpretations simultaneously. |
| Approach: | They propose a computational framework to study poorly-posedness of questions by generating spaces of potential interpretations and computing distributions based on interpretations chosen by answerers in the Reddit question thread. |
| Outcome: | The proposed framework analyzes poorly-posed questions using a set of interpretations chosen by human answerers and large language models. |
Language Modelling as a Multi-Task Problem (2021.eacl-main)
Copied to clipboard
| Challenge: | Using multitask learning, humans are optimising their behaviour towards a multitude of objectives to reach their goals in dayto-day life. |
| Approach: | They propose to study language modelling as a multi-task problem by examining the generalisation behaviour of language models as they learn the linguistic concept of Negative Polarity Items. |
| Outcome: | The proposed model is able to learn the linguistic concept of Negative Polarity Items (NPIs) and is a multi-task learning model. |
Penguins Don’t Fly: Reasoning about Generics through Instantiations and Exceptions (2023.eacl-main)
Copied to clipboard
| Challenge: | Generics express generalizations about the world that are not universally true . commonsense knowledge bases encode some generic knowledge but rarely enumerate exceptions . |
| Approach: | They propose a framework informed by linguistic theory to generate exemplars for generics . they generate 19k exemplar cases for 650 generics and show they outperform a strong baseline . |
| Outcome: | The proposed framework outperforms a baseline framework by 12.8 precision points. |
Variance of Average Surprisal: A Better Predictor for Quality of Grammar from Unsupervised PCFG Induction (P19-1)
Copied to clipboard
| Challenge: | In unsupervised grammar induction, data likelihood is only weakly correlated with parsing accuracy, especially at convergence after multiple runs. |
| Approach: | They propose to use VAS instead of data likelihood to find better grammars by examining linguistically-motivated constraints related to syntax. |
| Outcome: | The proposed model is better suited for word order typology classification than data likelihood. |
What company do words keep? Revisiting the distributional semantics of J.R. Firth & Zellig Harris (2022.naacl-main)
Copied to clipboard
| Challenge: | linguists J.R. Firth and Zellig Harris are often credited with the invention of "distributional semantics" a close reading of their work uncovers two distinct and in many ways divergent theories of meaning . |
| Approach: | They propose to compare two different theories of meaning that focus on internal workings of linguistic forms with a broader cultural and situational context. |
| Outcome: | The authors examine the differences between their theories of meaning and the internal workings of linguistic forms . they find that Firth could guide the field towards a more culturally grounded notion of semantics . |
CoPrUS: Consistency Preserving Utterance Synthesis towards more realistic benchmark dialogues (2025.coling-main)
Copied to clipboard
| Challenge: | Large-scale Wizard-Of-Oz dialogue datasets lack certain types of utterances, which would make them more realistic. |
| Approach: | They propose to use a large language model to create and repair communication errors in an automatic pipeline. |
| Outcome: | The proposed method is based on linguistic theory and uses a state-of-the-art Large Language Model (LLM) to create the error and repair it. |
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | a new analysis leverages linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs). |
| Approach: | They propose to use linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs). |
| Outcome: | The proposed analysis reveals that linguistic similarity is significantly influenced by training data exposure, leading to higher cross-LLM agreement in higher-resource languages. |
Language-specific Effects on Automatic Speech Recognition Errors for World Englishes (2022.coling-1)
Copied to clipboard
| Challenge: | Existing systems are not able to meet the needs of speakers of different demographic groups. |
| Approach: | They propose to analyze the performance of Otter’s automatic captioning system on native and non-native English speakers of different language background through a linguistic analysis of segment-level errors. |
| Outcome: | The proposed system predicts certain errors from the phonological structure of a speaker’s native language. |
Revisiting Supertagging for faster HPSG parsing (2024.emnlp-main)
Copied to clipboard
| Challenge: | a new supertagger for HPSG-based treebanks is used to improve parsing speed and accuracy. |
| Approach: | They propose to integrate the best supertagger into an HPSG-based parser and compare it to an existing system. |
| Outcome: | The proposed system achieves 97.26% accuracy on 950 sentences from WSJ23 and 93.88% on the out-of-domain technical essay The Cathedral and the Bazaar. |
Causal Interventions Reveal Shared Structure Across English Filler–Gap Constructions (2025.emnlp-main)
Copied to clipboard
| Challenge: | Language Models (LMs) have emerged as powerful sources of evidence for linguists seeking to develop theories of syntax. |
| Approach: | They propose to use causal interpretability methods to characterize abstract mechanisms that LMs learn to use by transferring a wh-filler-gap structure into a gap-less c++ class. |
| Outcome: | The proposed methods can characterize the abstract mechanisms that LMs learn to use, and challenge claims that they can be learned only with strong innate priors. |
Specifying Genericity through Inclusiveness and Abstractness Continuous Scales (2024.lrec-main)
Copied to clipboard
| Challenge: | Using a pilot study, we created a small but crucial annotated dataset of 324 sentences, demonstrating the framework’s effectiveness in capturing nuanced aspects of genericity. |
| Approach: | They propose a framework for fine-grained modeling of noun phrases' genericity in natural language using a small but crucial annotated dataset of 324 sentences. |
| Outcome: | The proposed framework can be used to model genericity of noun phrases in natural language and can be easily compared with existing binary annotations. |
Do Language Models Use Logophoric Cues? Evidence from Mandarin Chinese Long-Distance Reflexive (2026.findings-acl)
Copied to clipboard
| Challenge: | Using minimal pairs and surprisal-based measures, we assess whether large language models exhibit systematic biases toward non-local antecedents in logophoric contexts. |
| Approach: | They examine large language models’ sensitivity to four logophoric cues known to license long-distance binding of the reflexive ziji . |
| Outcome: | The proposed model families show that they exhibit above-chance sensitivity to all four cues, while lexically anchored cue are more robustly captured than discourse-level cue. |